Process phase
After creating a collection of files, use the Process phase to process and store the files to enable comprehensive search. While you move the collection through the Process phase, Epiq Discover performs the processing job. When a processing job is complete, the field displays the processed date and time for the documents except for excluded archives or file extensions.
Top level archive processing
For projects created in Epiq Discover release 4.6 or later, when processing an archived file, after extracting the top-level archive, the system creates a record for it and then deletes or moves its native file to DataHub based on the project settings. By default, the system moves the top‑level archive to DataHub. You can view the top‑level archive record in Inspect. The system populates metadata fields for this record. In DataHub, the system saves top‑level archives in a folder labeled - under > .
For the following file types, the system does not move the top-level archive to DataHub or delete it.
-
Encrypted files
-
Corrupt files
-
Partially extracted files
To configure the top-level archive file processing setting, refer to Configure top-level archive file processing setting.
Mid level archive processing
For projects created in Epiq Discover release 4.7 or later, the system sets the field of mid‑level archive files such as email archives and compressed or packaged formats to ExcludedArchive during processing when extraction is successful. As a result, the system excludes these files from promotion and export.
-
If the system does not successfully extract mid‑level archive files, it does not set the field to ExcludedArchive. Therefore, the system can promote or export these files.
-
The system does not process OneNote (.one) files as Excluded Archive.
Native email deduplication
During processing, the system automatically runs the Native Deduplication job. This job removes duplicate native email files.
Automatic OCR
During the Process phase, Epiq Discover automatically performs OCR on specific types of documents. The file types that can be automatically OCR'd include: JPG, TIF, TIFF, PDF, and Bitmap files. As an administrator, you can enable or disable this feature in Project Settings. Automatic OCR is enabled by default.
Automatic language detection
During the Process phase, Epiq Discover automatically detects languages in the documents that contain at least 250 characters and then populates the and fields.
-
For more information about the and fields, refer to Search field reference.
Some limitations apply to this feature. For more information, refer to Known limitations.
The following list provides related topics.